feat(prediction): audit an evaluation setup for leakage risks - #592
Open
breimanntools wants to merge 1 commit into
Open
breimanntools wants to merge 1 commit into
breimanntools wants to merge 1 commit into
Conversation
Leakage inflates a score without ever raising an error, and on small windowed datasets it is easy to introduce and hard to see. aa.audit_leakage runs a set of cheap heuristics over whatever parts of a setup are supplied and returns a plain DataFrame of findings (check, severity, detail, ids), worst first, with the overall verdict in df_audit.attrs["status"] so no bespoke result type is needed. Without splits it audits the dataset, which is the check to run before choosing a split; with splits it also compares the training against the test part within each fold. It reports duplicate sequences, a row index, protein or group present on both sides of a fold, a feature correlating with the label at |r| >= 0.95, an empty or strongly uneven fold, and a single-class or skewed fold. Windowed input is handled correctly: the sampler emits both the repeated parent 'sequence' and the 'window' cut from it, so the window takes precedence and sibling windows of one protein are not mistaken for duplicate samples. raise_on turns findings at or above a severity into a bare ValueError naming them; the default reports only and never raises, so an audit can be dropped into a workflow without changing its control flow. The checks are heuristic and a clean report is not a proof that no leakage exists; the severities are labels for a human reader, not a machine taxonomy, since whether a finding must block a workflow is a policy decision left to the caller. Pairs with bind_groups, which prevents the group leak this reports: a split from bind_groups(...).split(X, y) feeds straight into splits=. Refs #479 Co-Authored-By: Claude Fable 5.1 <[email protected]>
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #592 +/- ##
==========================================
- Coverage 95.39% 95.35% -0.04%
==========================================
Files 222 224 +2
Lines 23387 23710 +323
Branches 4073 4146 +73
==========================================
+ Hits 22309 22609 +300
- Misses 631 644 +13
- Partials 447 457 +10
... and 1 file with indirect coverage changes
🚀 New features to boost your workflow:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Most evaluation leaks are invisible until someone notices an implausibly high score.
aa.audit_leakageinspects a dataset and a split and returns a plain table of findings.Eight checks: duplicate sequences, train/test overlap, duplicates across folds, the same protein in several folds, group overlap, a target-derived feature, fold-size anomaly, class-balance anomaly. Dataset-level checks run without
splits; fold-level checks compare train against test within each fold.Returns a plain 4-column DataFrame (
check,severity,detail,ids), worst first, with the verdict indf.attrs["status"]— no bespoke result type, which the issue explicitly forbids. Pairs withbind_groups:splits=bind_groups(...).split(X, y)is the natural input.KPIs
Asserted both ways: a duplicate placed in two folds produces a high finding naming the ids; a clean setup returns an empty table with
status == "ok";raise_on="high"raises only when a high finding exists.Verification
64 tests (44 + 20 across the two house classes), hypothesis on 4, error-message
match=throughout. 1189 tests pass acrossprediction_tests+api_tests. pyright 0 errors. Docstring checker 0 defects, drift 0. Docs gate at baseline (167). Notebook: 7/7 public params by name, no markdown headings.Two things to look at
detail, rather than one row per fold — a 10-fold leak would otherwise emit 10 near-identical rows. That trades machine-addressable fold ids for readability; it is the one genuine design fork here.|r| >= 0.95, fold-size ratio 2.0, class-share deviation 0.2, ids capped at 10), so a report means the same thing everywhere. Documented inNotes.One test caught a true positive that is worth knowing:
GroupKFoldover proteins whose label is the protein's parity yields single-class folds, and the audit correctly reports it. The test expectation was fixed, not the code, and the notebook explains it.Refs #479 (no closing keyword — shipped code is a first draft).
🤖 Generated with Claude Code